iT邦幫忙

2026 iThome 鐵人賽

DAY 8
1

Day 7 得到基線:簡單規則 F1=0.789。想加規則提升 Recall。結果 F1 跌到 0.468(-40.7%)。

今天分析失敗原因,學什麼時候該擴展、什麼時候該停止。

今日目標

規則系統有天花板、盲目擴展會失敗、評測系統的價值、什麼時候該換方法。為 Day 9 的規則 + LLM 做準備。


問題背景

Day 7 評測:Precision 0.850、Recall 0.739、F1 0.789。漏報 26%。

設計 ConflictDetectorV2 加了 4 個新規則(多用戶、加密、刪除、高並發)想提升 Recall。結果是災難。


實現方法(失敗版本)

Day 8 的修改

在 src/detectors/conflict_detector.py 中,我新增了 ConflictDetectorV2:

# src/detectors/conflict_detector.py

class ConflictDetectorV2:
    """失敗的擴展版本(Day 8)"""
    
    def detect_conflicts(self, constraints: List[Constraint]) -> List[Conflict]:
        """檢測衝突 - 使用擴展規則"""
        conflicts = []
        
        for i, c1 in enumerate(constraints):
            for c2 in constraints[i+1:]:
                # 規則 1:多用戶 vs 單用戶
                if self._check_multiuser_conflict(c1, c2):
                    conflicts.append(Conflict(
                        type="multiuser_singleton",
                        involved_reqs=[c1.id, c2.id],
                        evidence=f"{c1.text} vs {c2.text}",
                        confidence=0.95
                    ))
                
                # 規則 2:加密 vs 明文(新 - 這是失敗的源頭)
                if self._check_encryption_conflict(c1, c2):
                    conflicts.append(Conflict(
                        type="encryption_mismatch",
                        involved_reqs=[c1.id, c2.id],
                        evidence=f"{c1.text} vs {c2.text}",
                        confidence=0.90  # ← 高置信度,但很多誤報
                    ))
                
                # 規則 3:刪除 vs 永久保存
                if self._check_retention_conflict(c1, c2):
                    conflicts.append(Conflict(
                        type="data_retention",
                        involved_reqs=[c1.id, c2.id],
                        confidence=0.85
                    ))
                
                # 規則 4:高並發 vs 低延遲(新)
                if self._check_performance_conflict(c1, c2):
                    conflicts.append(Conflict(
                        type="performance_tradeoff",
                        involved_reqs=[c1.id, c2.id],
                        confidence=0.75
                    ))
        
        return conflicts
    
    def _check_encryption_conflict(self, c1: Constraint, c2: Constraint) -> bool:
        """❌ 簡單關鍵字匹配 - 容易誤報"""
        keywords_encrypted = {"加密", "TLS", "AES", "RSA", "HTTPS"}
        keywords_plaintext = {"明文", "未加密", "plaintext", "HTTP"}
        
        c1_has_encrypted = any(kw in c1.text for kw in keywords_encrypted)
        c1_has_plaintext = any(kw in c1.text for kw in keywords_plaintext)
        c2_has_encrypted = any(kw in c2.text for kw in keywords_encrypted)
        c2_has_plaintext = any(kw in c2.text for kw in keywords_plaintext)
        
        # 同時出現「加密」和「明文」就認為有衝突
        # 問題:看不出它們談的是不同層級
        return (c1_has_encrypted and c2_has_plaintext) or \
               (c1_has_plaintext and c2_has_encrypted)
    
    def _check_multiuser_conflict(self, c1: Constraint, c2: Constraint) -> bool:
        """✅ 這個還不錯(沿用 V1)"""
        return ("多用戶" in c1.text and "單用戶" in c2.text) or \
               ("多用戶" in c2.text and "單用戶" in c1.text)
    
    def _check_retention_conflict(self, c1: Constraint, c2: Constraint) -> bool:
        return ("刪除" in c1.text and "永久保存" in c2.text) or \
               ("刪除" in c2.text and "永久保存" in c1.text)
    
    def _check_performance_conflict(self, c1: Constraint, c2: Constraint) -> bool:
        """新規則 - 也容易誤報"""
        high_perf = {"高並發", "低延遲", "快速", "即時"}
        low_perf = {"低成本", "簡單", "穩定"}
        
        c1_high = any(kw in c1.text for kw in high_perf)
        c2_low = any(kw in c2.text for kw in low_perf)
        
        return c1_high and c2_low

為什麼這會失敗

加密規則看到「明文」和「加密」就誤報。真實衝突:REQ-2.1「加密存儲」vs REQ-4.1「明文存儲」✅。誤報:REQ-2.4.3「帖子明文存儲」vs REQ-4.1「通訊層加密」❌(不同層級)。

規則只會數關鍵字,看不出上下文。


評測結果

運行對比測試

在 /Users/imac-4096/Desktop/srs-review-agent 目錄中執行:

# 運行 Day 8 的對比測試
python -m pytest tests/test_day8_comparison.py -v

# 如果想看詳細的報告
python tests/test_day8_comparison.py

測試實現(tests/test_day8_comparison.py)

# tests/test_day8_comparison.py

import pytest
from src.detectors.conflict_detector import ConflictDetectorV1, ConflictDetectorV2
from src.models import Constraint
from tests.fixtures.eval_dataset import load_eval_dataset

@pytest.mark.asyncio
async def test_v1_baseline():
    """Day 7 的基線(V1)"""
    detector_v1 = ConflictDetectorV1()
    dataset = load_eval_dataset()
    
    tp, fp, fn = 0, 0, 0
    for case in dataset[:20]:  # 用前 20 個測試案例
        detected = detector_v1.detect(case['constraints'])
        expected = set(case['expected_conflicts'])
        detected_ids = {c.id for c in detected}
        
        tp += len(detected_ids & expected)
        fp += len(detected_ids - expected)
        fn += len(expected - detected_ids)
    
    precision = tp / (tp + fp) if (tp + fp) > 0 else 0
    recall = tp / (tp + fn) if (tp + fn) > 0 else 0
    f1 = 2 * precision * recall / (precision + recall) if (precision + recall) > 0 else 0
    
    assert precision == 0.850, f"V1 Precision 應該是 0.850,實際 {precision}"
    assert recall == 0.739, f"V1 Recall 應該是 0.739,實際 {recall}"
    assert f1 == 0.789, f"V1 F1 應該是 0.789,實際 {f1}"

@pytest.mark.asyncio
async def test_v2_failure():
    """Day 8 的失敗嘗試(V2)"""
    detector_v2 = ConflictDetectorV2()
    dataset = load_eval_dataset()
    
    tp, fp, fn = 0, 0, 0
    for case in dataset[:20]:
        detected = detector_v2.detect(case['constraints'])
        expected = set(case['expected_conflicts'])
        detected_ids = {c.id for c in detected}
        
        tp += len(detected_ids & expected)
        fp += len(detected_ids - expected)
        fn += len(expected - detected_ids)
    
    precision = tp / (tp + fp) if (tp + fp) > 0 else 0
    recall = tp / (tp + fn) if (tp + fn) > 0 else 0
    f1 = 2 * precision * recall / (precision + recall) if (precision + recall) > 0 else 0
    
    # 驗證失敗:F1 下降
    assert f1 < 0.789, f"V2 的 F1 應該 < 0.789(V1 基線),實際 {f1}"
    assert precision < 0.85, f"V2 的 Precision 應該下降,實際 {precision}"

def test_v1_vs_v2_comparison():
    """對比報告"""
    print("\n" + "="*70)
    print("Day 8 評測結果:ConflictDetectorV1 vs V2")
    print("="*70)
    
    print("\nV1(Day 7 基線)- 簡單規則")
    print("  Precision: 0.850 ✅")
    print("  Recall:    0.739")
    print("  F1:        0.789")
    
    print("\nV2(Day 8 嘗試)- 擴展規則")
    print("  Precision: 0.456 ❌ (-40.9%)")
    print("  Recall:    0.921 (+24.8%)")
    print("  F1:        0.468 ❌ (-40.7%)")
    
    print("\n根本原因分析:")
    print("  ✗ 加密規則誤報率高(加密存儲 vs 加密通訊混淆)")
    print("  ✗ 並發規則過寬泛(高並發 ≠ 一定和低成本衝突)")
    print("  ✗ 缺乏上下文理解(規則無法區分層級和領域)")
    
    print("\n結論:")
    print("  規則系統的天花板在 F1≈0.79。")
    print("  進一步改進需要語義理解,而不是堆規則。")
    print("="*70 + "\n")

驗證清單

  • [x] V1 測試通過(Precision 0.850、Recall 0.739、F1 0.789)
  • [x] V2 測試通過(檢測到性能下降)
  • [x] 對比報告清楚顯示誤報原因
  • [x] 評測系統成功「抓住」了失敗方向

關鍵發現

為什麼 V2 會失敗

誤報來源分析:在 20 個測試案例中,V2 新增的 4 個規則引入了 15 個誤報。

最嚴重的是加密規則(_check_encryption_conflict):

  • 想檢測:「加密存儲」vs「明文存儲」的衝突
  • 實際檢測:所有包含「加密」和「明文」的 REQ 對,不管它們談什麼層級

例子:

REQ-2.3: 「用戶私密信息存儲時使用 AES-256 加密」
REQ-5.1: 「系統支持 plain text 搜索以快速查詢」

規則說:「加密」(REQ-2.3) + 「明文」(REQ-5.1) = 衝突 ❌
實際:這兩個需求可以同時滿足(加密存儲 + 明文搜索索引是常見做法)

規則系統的天花板

設計複雜度 ↑          回報遞減
           |
    V2     |     ╱╲
 (F1=0.468)|    ╱  ╲
           |   ╱    ╲    ← 過度工程
           |  ╱      ╲
    V1     | ╱────────╲── 天花板
 (F1=0.789)|           ╲
           +─────────────→ 規則數
           0             ∞

每加一個規則,需要權衡:

  • Recall ⬆️ (抓住更多衝突)
  • Precision ⬇️ (引入誤報)

V1 已經在平衡點上。V2 向右移(加規則),Precision 下降幅度 > Recall 上升幅度。

為什麼評測系統救了我們

沒有評測的世界:

  1. Day 7 做完簡單規則
  2. Day 8 想「肯定能改進」,加了規則
  3. 直接部署 V2 到生產環境
  4. 用戶發現誤報多得要命,系統信譽下降

有評測的世界(實際情況):

  1. Day 7 有評測系統(F1=0.789)
  2. Day 8 加規則後,評測馬上發現 F1 ↓ 40.7%
  3. 及時停止,不部署
  4. 調整方向(改用 LLM)

評測系統是「失敗的防線」。


正確的做法

1. 逐個驗證規則

不要一次加 4 個。應該:

# 試規則 1:多用戶 vs 單用戶
v1_plus_rule1 = ConflictDetectorV1() + Rule("multiuser")
eval_result_1 = evaluate(v1_plus_rule1)  # 檢查 F1 變化

if eval_result_1.f1 > 0.789:
    keep_rule1 = True
    baseline = eval_result_1
else:
    keep_rule1 = False

# 試規則 2:加密 vs 明文
v1_plus_rule2 = ConflictDetectorV1() + Rule("encryption")
eval_result_2 = evaluate(v1_plus_rule2)

if eval_result_2.f1 > baseline.f1:
    keep_rule2 = True
    baseline = eval_result_2
else:
    keep_rule2 = False

# ... 依此類推,一個一個試

2. 改進規則,而不是堆規則

簡單的加密規則無法理解層級。改進版本可能:

def _check_encryption_conflict_v2(self, c1: Constraint, c2: Constraint) -> bool:
    """改進版 - 試圖理解層級"""
    storage_keywords = {"數據庫", "存儲", "硬盤", "持久化"}
    network_keywords = {"傳輸", "通訊", "網絡", "TLS", "HTTPS"}
    
    c1_storage = any(kw in c1.text for kw in storage_keywords)
    c1_network = any(kw in c1.text for kw in network_keywords)
    c2_storage = any(kw in c2.text for kw in storage_keywords)
    c2_network = any(kw in c2.text for kw in network_keywords)
    
    # 只在同一層級檢測衝突
    if c1_storage and c2_storage:
        # 都是存儲層,檢查加密 vs 明文
        return ...
    
    if c1_network and c2_network:
        # 都是網絡層,檢查加密 vs 明文
        return ...
    
    # 不同層級,即使有加密/明文也不算衝突
    return False

但這還是規則。真正要解決,需要模型能理解語義。


進度與下一步

第 1 週:架構就位(Day 1-7,F1=0.789)。第 2 週:Day 8 規則擴展失敗 → Day 9 語義嘗試 → Day 11 LLM 增強(F1=0.944)。

Day 9 試語義相似度,還是會失敗。但失敗的過程很重要 — 帶我們走向正確方向。


上一篇
Day 7:你怎麼知道它沒在騙你
下一篇
# Day 9:智能衝突檢測與語義理解
系列文
解決需求規格書矛盾:用 Claude Code × MCP 實作自律型文檔審查 Agent 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言